Back

The Annals of Applied Statistics

Institute of Mathematical Statistics

Preprints posted in the last 7 days, ranked by how well they match The Annals of Applied Statistics's content profile, based on 19 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Statistical Inference and Power Analysis for Comparative F1 and Fβ Scores under Correlated Classifier Pairs

Hsu, C.-Y.; Liu, Q.; Shyr, Y.

2026-07-17 dermatology 10.64898/2026.07.15.26358166 medRxiv
Top 0.2%
0.8%
Show abstract

As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.

2
The Variance-Stabilizing Transformation for the Poisson Rate Ratio: Closed-Form Confidence Intervals

Ng, S.-P.

2026-07-18 epidemiology 10.64898/2026.07.16.26358255 medRxiv
Top 0.5%
0.3%
Show abstract

The incidence rate ratio R is the standard measure for comparing event rates in clinical trials and epidemiology. In vaccine trials, the vaccine efficacy is VE = 1 - R. When events are rare, the two arm counts are Poisson. The estimator of R is heteroskedastic: its sampling variance changes with the data. So no fixed-width interval covers correctly everywhere. The usual log-Wald interval is undefined at zero events and covers poorly at small counts. Early vaccine and drug-safety readouts fall in exactly this regime. We show that a single reparameterization collapses this bivariate problem to an effective one-parameter family with a quadratic variance function, whose variance-stabilizing transformation is 2 arcsinh(sqrt(R)). The reduction yields a closed-form confidence interval for R. Its two leading errors, a curvature bias and the variability of the estimated scale, each admit a closed-form correction with no tuning constants. In a Monte Carlo study of our seven arcsinh variants and five competitors, the +Curve+Stu variant covers within 0.002 of the nominal 0.95 for about 50 control and 5 treatment events. Its width is on par with the best competitor. It avoids the conservatism and zero-count breakdown of log-Wald and MOVER. For moderate counts, we recommend this interval; for sparser data, our Bar-Lev and Enis count-shift variant is more robust. The result is a ready-to-use, closed-form interval for the low-count regime. We illustrate it on early Covid-19 vaccine-efficacy readouts and provide reference implementations in R and Python.

3
Genetic sensitivity analysis: estimating genetic confounding and environmentally mediated genetic effects using multiple exposures

Frach, L.; Rijsdijk, F.; Hannigan, L. J.; Dudbridge, F.; Pingault, J.-B.

2026-07-17 epidemiology 10.64898/2026.07.16.26358236 medRxiv
Top 0.7%
0.2%
Show abstract

Polygenic scores are imperfect measures of the additive genetic effects of common genetic variants. The resulting measurement error biases estimates of quantities of interest in epidemiological analyses integrating polygenic scores. For example, how much of an exposure-outcome association is genetically confounded can be substantially underestimated when using polygenic scores alone. Here we present extensions to Gsens, a genetic sensitivity analysis, which aims to correct for such measurement error using both polygenic scores and heritability estimates. Gsens now allows for multiple exposures and estimates several quantities of interest, i.e. genetic confounding, adjusted residual association (net of genetic confounding), genetic overlap and environmentally mediated genetic effects. We present derivations and simulations showing how Gsens accounts for measurement error in the polygenic score; we also show how estimation may be affected by misspecifications of the causal structure between exposures. Applying Gsens in the Norwegian Mother, Father and Child Cohort Study (MoBa), we uncover, among other results, substantial genetic confounding in the associations between multiple known risk factors for attention deficit hyperactivity disorder (ADHD), such as low birth weight and temperament, and measures of ADHD in childhood. The updated Gsens R package offers multiple options, including for missing data handling and customisable syntax. Our extended version of Gsens is applicable to a broad range of substantive questions in multiple disciplines.

4
Dual-Filament 3D Printing of Patient-Specific CT Phantoms with Embedded Implants and Tunable Metal-Artifact Intensity

Pasyar, P.; Mei, K.; Im, J. Y.; Roshkovan, L.; Geagan, M.; Noël, P. B.

2026-07-20 radiology and imaging 10.64898/2026.07.17.26358319 medRxiv
Top 0.8%
0.1%
Show abstract

ABSTRACT Background: Metallic implants such as orthopedic screws, prostheses, and dental hardware produce beam-hardening, photon-starvation, and streak artifacts that degrade computed tomography (CT) image quality, and the metal artifact reduction (MAR) methods developed to mitigate them require objective, reproducible benchmarking. Purpose: Objective evaluation of MAR algorithms in CT is hindered by the absence of phantoms that simultaneously provide anatomically realistic backgrounds, embedded implants of known geometry, and controllable, ground-truth--referenced artifact intensity. We present a dual-filament, voxel-level three-dimensional (3D) printing method that fulfills these requirements and demonstrate its capabilities on a clinically representative cervical spine case with embedded orthopedic spinal screws. Methods: The proposed method extends the PixelPrint framework, a fused-deposition-modeling (FDM) workflow that converts clinical Digital Imaging and Communications in Medicine (DICOM) data directly into 3D-printer Geometric code (G-code) without intermediate segmentation or surface meshing, to interleaved, voxel-level deposition of two filaments: a calcium-doped polylactic acid (PLA) for soft tissue and bone, and a higher-attenuation metal-doped PLA for metallic implants. For demonstration, anonymized DICOM data of a healthy cervical spine were used to design and fabricate three matched phantoms, each with six embedded spinal screws at C4--C6: a 0% metal-infill ground-truth phantom, a 50% medium-metal-infill phantom, and an 85% high-metal-infill phantom. All phantoms were scanned on a clinical spectral CT system at 120 kVp and 1000 mAs, reconstructed at 0.67 mm slice thickness with virtual monoenergetic imaging (VMI) across 50--190 keV. Method performance was characterized by region of interest (ROI)-based Hounsfield Unit (HU) agreement with the source patient data and by the noise-independent Gumbel-distribution p-index metric. Results: The dual-filament method reproduced patient anatomy, soft-tissue contrast, and screw geometry with high fidelity. ROI HU values agreed with patient data within {+/-}25 HU for soft tissue and trabecular bone; cortical regions were underestimated owing to the current ceiling of the calcium-doped PLA used in this study. The tunable-artifact behavior was quantified as follows: the Gumbel location parameter scaled monotonically from 46.7 HU (no-metal background) to 57.1 HU (50% infill) to 90.5 HU (85% infill) for the VMI 70 keV with standard filter. High-keV VMI reconstructions substantially reduced streak and beam-hardening artifacts while preserving anatomic detail. Conclusions: The proposed dual-filament, voxel-level PixelPrint method enables the fabrication of patient-specific, multi-material CT phantoms with embedded metallic implants and controllable, ground-truth--referenced artifact intensity. Although demonstrated here in a single cervical-spine case, the workflow is anatomy- and implant-agnostic by construction and could in principle be adapted to other musculoskeletal sites (e.g., knee, hip, dental) and implant materials, providing a reproducible methodological foundation for benchmarking MAR algorithms, characterizing spectral CT performance, and validating emerging photon-counting detector systems. Keywords: 3D printing methodology; fused deposition modeling; voxel-level multi-material printing; spectral computed tomography; metal artifact reduction; phantom design; orthopedic implants; dual filament; PixelPrint.

5
Development and external validation of deep learning models for spontaneous preterm birth prediction from mid-trimester cervical ultrasound

Chanian, R.; Mishra, D.; Jain, R.; Sharma, N.; Khurana, A.; Tripathi, R.; Tripathi, A.; group, G.-I. s.; Wadhwa, N.; Noble, J. A.; Thiruvengadam, R.; Desiraju, B. K.; Bhatnagar, S.

2026-07-19 obstetrics and gynecology 10.64898/2026.07.17.26358221 medRxiv
Top 1%
0.1%
Show abstract

Preterm birth is the leading cause of neonatal death. Despite sustained efforts to identify high-risk women in the mid-trimester, accurate prediction remains difficult. Quantitative cervical ultrasound texture has been proposed as a predictor of spontaneous preterm birth. However, earlier models were developed in small single-centre samples and were not externally validated. We developed image-texture (Local Binary Patterns with a Random Forest), deep-learning (Vision Transformer), clinical-variable, and multimodal models to predict spontaneous preterm birth on the prospective GARBH-Ini cohort. We then externally validated our best models on an independent cohort scanned on a different ultrasound machine. Our best overall model reached an internal-test area under the receiver-operating-characteristic curve of 0.71 (95% CI 0.60, 0.82), but performed modestly at 0.52 (95% CI 0.38, 0.64) externally. The deep-learning and multimodal models did not perform better. Discrimination appeared higher in a clinically high-risk subgroup at the 34-week threshold. These estimates were imprecise because of few cases and need to be confirmed in future studies. Among the several likely reasons for the modest external performance is the heterogeneity of preterm birth. Predicting distinct preterm-birth subtypes separately, and integrating additional biomarkers and data domains, might improve model performance. Keywords: preterm birth; cervical ultrasound; prediction model; external validation; deep learning

6
Efficient stochastic epidemic simulation via the Sellke construction

van Boven, M.; Bootsma, M. C.

2026-07-17 epidemiology 10.64898/2026.07.16.26358219 medRxiv
Top 1%
0.1%
Show abstract

Stochastic epidemic models are a cornerstone of infectious disease epidemiology and are often used to study intervention scenarios. However, large run-to-run variability can make intervention effects difficult to estimate precisely. We revisit the epidemic Sellke construction, which assigns each individual an infection threshold for the cumulative infection hazard such that, conditional on the thresholds, the epidemic trajectory becomes deterministic. This enables coupling of simulations with and without an intervention, yielding low-variance effect estimates even when outcomes such as final size or peak incidence vary widely between runs. We develop an exact, event-driven implementation that maintains infection and recovery events in priority queues. Cumulative infection-hazard updates require O(log N) time per event, yielding overall complexity O(Elog N) for E events in a population of size N. The implementation achieves computational performance comparable to the classical Gillespie algorithm while naturally accommodating non-Markovian infectious periods and complex infectiousness profiles. We illustrate the approach using distance-dependent spread of avian influenza between poultry farms in the Netherlands and a multilayer population with households, schools, and workplaces. In both examples, coupling enables efficient within-run comparisons of intervention scenarios across stochastic realisations.

7
A Study Of Factors Influencing Fetal Monitor Failure

Tsanligrenchin, D.; Enkhjargal, E.-U.; Boldbaatar, O.; Shagdar, I.; Tumurtogoo, A.; Tuya, A.; Batbold, S.

2026-07-15 obstetrics and gynecology 10.64898/2026.07.12.26357884 medRxiv
Top 1%
0.1%
Show abstract

In Mongolia, an average of 65,000 women become pregnant each year, and about 59,500 babies are born. Although the number of pregnancies is decreasing by 8-12 percent each year, the level of fetal monitor usage remains high. The capital's maternity hospital currently has 27 fetal monitors in use, and an average of 30-35 calls are recorded per month. However, there is a lack of research on the use of fetal monitors, the causes and influencing factors of damage, and the organization of technical services. Therefore, this topic was chosen to determine the usage status of fetal monitors, the causes of malfunctions, and ways to improve them. Purpose To study the causes and factors affecting possible damage and injury during the use of fetal monitors, and to identify ways to reduce them. Materials and methods A one-time study was conducted on 10 MT-610 fetal monitors that were put into operation in 2019 at the Urgo Maternity Hospital in the capital. Data were collected and processed using document analysis methods from the technical passports and call logs of these devices. The factors contributing to common failures were identified using focus group interviews with the engineers and technicians responsible for the equipment. Results This study found that fetal monitor failures are caused by improper use, lack of regular calibration, electrical fluctuations, ambient temperature and humidity, and insufficient medical staff skills, training, and knowledge of how to use the device, all of which contribute to failures and measurement errors. It is also observed that when a replacement part is needed for a monitor that frequently breaks, the monitor is more likely to break again if it is used as a replacement from a previously broken monitor. Therefore, training doctors and nurses who replace spare parts on their use has been observed to significantly reduce future breakdowns. Conclusion According to the study results, the breakdowns and failures of fetal monitoring devices are mainly related to internal system failures, unstable power supply, and wear and tear of accessories and mechanical parts. The highest percentage of device failures indicates the need for special attention to the reliability of the device's basic functions. Additionally, the high percentage of accessory and printer failures indicates the need for proper use and monitoring of the entire device. In addition to technical factors, human misuse, lack of maintenance, and environmental influences also play a significant role in damage. Therefore, it is concluded that to ensure the reliable operation of fetal monitors, it is necessary to perform regular maintenance, stabilize the power supply, improve the quality of accessories, and increase the knowledge and skills of medical staff. Keywords: Fetal Monitoring, Equipment Failure, Risk Factors

8
Number of previous caesarean deliveries, history of vaginal birth and uterine rupture during a trial of labour : Protocol for a retrospective population-based cohort study

Thompson, R.; De Vries, B.; Adily, P.; Narayan, R.; Mackie, A.; Phipps, H.; Berghella, V.; Lauer, M.

2026-07-21 obstetrics and gynecology 10.64898/2026.07.19.26358407 medRxiv
Top 1%
0.1%
Show abstract

There remains considerable uncertainty around the safety of a trial of labour after more than one caesarean delivery. This large retrospective cohort study will investigate the safety of a trial of labour using a large United States dataset of all registered births from 2011 to the most recent year with available data. A multivariable fitted model will be used to predict the probability of uterine rupture in women with two or more previous caesarean deliveries, with a minimum 21-month interpregnancy interval. This will provide important information to clinicians in counselling women wanting to attempt a vaginal birth after more than one caesarean delivery.

9
From Lotka-Volterra Dynamics to Community Assembly: Theory, Topography, and Empirical Applications

Schreiber, S.; Brennan, J.; Spaak, J. W.

2026-07-15 ecology 10.64898/2026.07.14.738515 medRxiv
Top 1%
0.1%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWO_LICommunity assembly graphs (CAGs) summarize which species combinations can coexist and how single-species invasions drive transitions between them, encoding the pathways, alternative endpoints, and cycles that make up a communitys assembly history. Constructing CAGs from dynamical models requires methods that are both computationally tractable and faithful to the underlying ecological dynamics. However, existing methods rely on restrictive assumptions, such as global stability, that exclude alternative stable states and non-equilibrium dynamics known to occur in empirical systems. C_LIO_LIWe develop a computational pipeline that constructs CAGs from any generalized Lotka-Volterra model. Building on the invasion graph framework and its connection to permanence, the pipeline verifies that community dynamics are bounded, identifies which subsets of species coexist in the sense of permanence, determines which single-species invasions are dynamically realized, and assigns each community a topographic height equal to the length of the longest assembly path leading to it. We also provide a numerical algorithm to simulate the dynamics of community assembly. C_LIO_LIWe prove several general properties of the resulting graphs, including that a successful invader is never subsequently excluded and that, in the absence of assembly cycles, permanent communities can be reassembled by introducing their species one at a time in the right order. We prove that the CAG faithfully reproduces the compositional shifts seen in the numerically simulated dynamics of assembly. Applying the pipeline to three empirically based models (a New Zealand grassland, a European pasture, and a Puerto Rican ant community), we show how competition strength and mutualistic feedbacks reshape the assembly landscape and how intransitive competition generates assembly cycles. C_LIO_LIOur approach accommodates alternative stable states and non-equilibrium dynamics without requiring global stability, and it turns the long-standing landscape metaphor into a quantitative, mechanistically grounded object by resolving what "height" means. More broadly, it makes the topography of the assembly pathways measurable, providing a way to compare the historical contingency and predictability of the assembly in ecological systems. C_LI

10
Where species distribution models fail under occurrence-data contamination: calibration error concentrates at stream-network headwaters

Miok, K.; Laza, A. V.; Skrlj, B.; Robnik-Sikonja, M.; Parvulescu, L.

2026-07-15 ecology 10.64898/2026.07.14.738364 medRxiv
Top 1%
0.1%
Show abstract

Species distribution models (SDMs) increasingly inform conservation and biosecurity decisions in freshwater systems, where the reliability of its uncertainty estimates matters as much as its point predictions. Ensemble SDMs derive prediction intervals from across-replicate variance, but this variance captures systematic error only when replicates disagree about it, an assumption that fails when training data are contaminated with low-accuracy records, the norm in citizen-science datasets. Whether this failure is spatially uniform or concentrates in identifiable parts of a range is unknown. Using a panel of European freshwater crayfish spanning native headwater-associated species and invasive lowland colonizers, we show that contamination-induced calibration failure is strongly spatially structured: it concentrates at stream-network headwaters, the topological tops of the network, where upstream-aggregated predictors are structurally undefined, and scales with contamination severity, replicated across four species and both dominant ensemble protocols (replicate and consensus). The failure is driven by upward prediction bias, not by intervals failing to widen: contaminated ensembles overpredict suitability in headwaters, and because the bias is shared across ensemble members, the intervals do not flag it. This is a conservation-relevant blind spot, because headwaters are both refugia for threatened native crayfish and front lines for invasion; an SDM that silently overpredicts suitability there misdirects survey and management effort toward the segments where its predictions are least trustworthy. Standard leave-one-basin-out conformal calibration, the recommended panel-wide remedy, repairs marginal coverage but leaves headwaters undercovered, because a single calibration threshold is dominated by the abundant non-headwater segments. A group-conditional (Mondrian) variant, calibrating the two populations separately, restores reliable coverage in both at no extra cost and reallocates width where it is needed. We recommend network-position-stratified calibration as a default for ensemble SDMs in dendritic freshwater systems.

11
DELENDA: Differentiable Epidemiology for Latent-state Estimation and Nonlinear Decision Analysis

Suresh, J.

2026-07-15 epidemiology 10.64898/2026.07.13.26357962 medRxiv
Top 1%
0.1%
Show abstract

Malaria subnational tailoring is often a population-level allocation problem: which interventions should be prioritized, at what coverage, and under what budget and uncertainty assumptions? We present DELENDA, a differentiable compartmental model of Plasmodium falciparum transmission designed for posterior calibration and intervention-mix optimization. We fit a NUTS posterior jointly to age-stratified prevalence and clinical-incidence data from five sub-Saharan African sites plus three pre-intervention Garki Project villages, spanning a broad entomological inoculation rate (EIR) range. DELENDA is implemented in JAX, which makes the full simulation differentiable. This enables efficient Bayesian inference and continuous constrained optimization over intervention coverage. We apply the framework to an illustrative decision problem: a highly seasonal transmission setting where coverage is optimized for ITNs, SMC, IRS, and pediatric malaria vaccination across EIR, budget, objective, and uncertainty grids. Three findings are decision-relevant. First, intervention rankings are more robust than projected impact: posterior, vector-biology, and intervention-efficacy uncertainty change optimized coverage modestly but substantially widen the distribution of cases averted. Second, the objective matters: under-five optimization brings child-targeted SMC and vaccination in earlier, whereas all-age optimization delays vaccination and favors broader population protection through IRS. Third, cost uncertainty is mainly a constraint-side problem: expected-cost optima have material budget-overrun probability, while tail-risk budget rules sharply reduce overrun risk at the cost of lower effective coverage and fewer expected cases averted. DELENDA therefore demonstrates an uncertainty-first approach to subnational tailoring: differentiable model structure exposes the biological parameter space to posterior calibration and carries biological and operational uncertainty into constrained decision optimization, tasks that are difficult with the non-differentiable models currently central to SNT workflows.

12
From Christmas sex to winter intimacy: three decades of birth seasonality, sex ratio dynamics, and fertility change in South Africa, 1994-2024

Masukume, R.; Chimberengwa, P. T.; Masukume, G.; Liczbinska, G.; Grech, V.; Mapanga, W.

2026-07-15 obstetrics and gynecology 10.64898/2026.07.12.26357853 medRxiv
Top 2%
0.1%
Show abstract

BACKGROUND: Since South Africa's democratic transition in 1994, the country has undergone profound social, demographic and public health change. We analysed national recorded live-birth data from 1994-2024 to identify major signals of population reproduction. METHODS: Monthly recorded live births from January 1994 to December 2024 were obtained from Statistics South Africa. Birth seasonality, sex ratio at birth (SRB) [male/total live births] and annual recorded live births were analysed using time-series and forecasting methods. RESULTS: From 1994-2014, September was the peak birth month in all 21 years, consistent with conceptions during the Christmas-New Year holiday period nine months earlier. From 2015 onwards, March became the most frequent peak month, with April emerging as the peak month in 2024, indicating a shift towards winter conceptions. The SRB declined to 49.996% in June 2021 (95% prediction interval 50.165%-50.749%) and remained below the lower prediction bound from May to July 2021 (combined p<0.001). November 2021 recorded the highest monthly SRB in the 31-year study period (50.983%), exceeding the upper 95% prediction interval. Annual recorded live births peaked at 1,112,378 in 2008 and declined to 798,556 in 2024; births from 2022-2024 fell below the 95% confidence interval of the historical trend. CONCLUSIONS: Three prominent demographic signals emerged: a shift from Christmas holiday conceptions towards winter conceptions; a rare inversion (SRB <50%) and sustained depression of the SRB during May-July 2021, occurring within the 3-5-month stress-sensitive window after the January 2021 Beta-wave mortality peak, followed by the highest monthly SRB of the study period in November, nine months after the easing of COVID-19 restrictions in February 2021; and an accelerated decline in annual recorded live births after 2021, culminating in the lowest level observed in 2024. These findings indicate changes in reproductive timing, stress-sensitive sex-ratio patterning and fertility in South Africa.

13
Comparing different neuroimaging modalities for quantification of the cholinergic system in Parkinson's disease

d'Angremont, E.; Marschall, T. M.; Renken, R. J.; Sommer, I. E.

2026-07-17 neurology 10.64898/2026.07.15.26357522 medRxiv
Top 2%
0.0%
Show abstract

Introduction Parkinson's disease (PD) is a multifactorial disorder, affecting multiple neurotransmitter systems, including the cholinergic system. Cholinergic denervation is heterogeneous across patients and difficult to predict based on clinical presentation. In this study, we assessed the sensitivity of structural MRI (sMRI) and functional MRI (fMRI) to cholinergic degeneration related to PD and to cognitive functioning in PD. We compared our results to results from previously reported [18F]Fluoroethoxybenzovesamicol ([18F]FEOBV) PET imaging, which is considered the gold standard for cholinergic imaging. Methods 34 PD patients and 10 healthy controls underwent structural T1-weighted MRI. A subset of 14 patients and 9 controls also underwent resting-state fMRI. We extracted the bilateral volumes of the nucleus basalis of Meynert (NBM) from the sMRI images. Functional connectivity (FC) from the NBM to the cortex (NBM-FC) was determined using fMRI data. Principal component analysis (PCA) was applied to reduce the dimensionality of the NBM-FC images. We assessed performances for NBM-FC in distinguishing patients from controls using stepwise logistic regression. Similarly, NBM volume was used using logistic regression. Furthermore, the relation between these measures and cognitive function in several domains was investigated with (stepwise) linear regression. Leave-one-out cross validation (LOOCV) and bootstrapping was performed to assess robustness of the results. Results NBM-FC was well able to discriminate patients from controls with an AUC of 0.84 (95% CI: 0.62-1). NBM volume showed lower performance, but was still better than chance: AUC: 0.75 (95% CI: 0.57-0.93). Significant correlations were found between 1) cognition in the attentional domain and NBM-FC (r=0.63; p=.015) and 2) global cognition and NBM volume (r=0.55, p=.001). These results were inferior to those previously reported using [18F]FEOBV tracer uptake (see Chapter 6). Bootstrapping revealed that NBM volume of only the left hemisphere was stably related to PD diagnosis and global cognition in PD patients. We found that a lower NBM-FC in specific brain areas, including the fusiform gyrus, supramarginal gyrus and dorsolateral prefrontal cortex, was related to PD diagnosis. Bootstrapping revealed no stable NBM-FC pattern related to attention. Conclusion Although MRI results were slightly inferior to [18F]FEOBV PET data, MRI may provide a cheaper and more widely available alternative for cholinergic imaging. We recommend testing the utility of MRI as predictor and monitor of cholinergic treatment effect in a longitudinal study.

14
An ancestry-matched Mendelian randomisation analysis of kidney function and heart failure subtypes in African ancestry populations

Gaye, N. D.; Diawara, A.

2026-07-17 genetic and genomic medicine 10.64898/2026.07.15.26358145 medRxiv
Top 2%
0.0%
Show abstract

Chronic kidney disease and heart failure disproportionately burden populations of African ancestry, yet Mendelian randomisation (MR) studies of the causal relationship between kidney function and heart failure subtypes have been conducted exclusively in European ancestry populations. We performed a forward two-sample MR analysis to evaluate the causal effect of genetically predicted estimated glomerular filtration rate (eGFR) on heart failure with preserved ejection fraction (HFpEF) and heart failure with reduced ejection fraction (HFrEF) in individuals of African ancestry. Genetic instruments were selected from an African ancestry eGFR genome-wide association study (N = 67,943) at genome-wide significance, with linkage disequilibrium clumping using an African ancestry reference panel. Heart failure subtype summary statistics were obtained from the Million Veteran Program (HFpEF: 5,379 cases / 113,041 controls; HFrEF: 9,104 cases / 109,632 controls). Six independent SNPs (F-statistics 30.5 &#8211 107.3; R&#178 = 0.62%) were retained as instruments. The primary inverse-variance weighted analysis provided no evidence of a causal effect of eGFR on HFpEF (OR 0.92, 95% CI 0.80 &#8211 1.06, p = 0.248) or HFrEF (OR 0.98, 95% CI 0.78 &#8211 1.23, p = 0.878). Sensitivity analyses were directionally consistent. There was no evidence of heterogeneity or directional pleiotropy. Minimum detectable effects at 80% power were OR 1.28 for HFpEF and OR 1.22 for HFrEF. These null findings should be interpreted as inconclusive given current power constraints; larger ancestry-matched studies are needed.

15
Portable Ultra-Low Field MRI Deep-Learning Algorithms for White Matter Lesion Segmentation Improve Accuracy and Reflect Clinical Disability in Multiple Sclerosis

Thommana, A. A.; Donnay, C. A.; Norato, G.; Gaitan, M. I.; Griffanti, L.; Nair, G.; Reich, D. S.; Okar, S. V.

2026-07-17 neurology 10.64898/2026.07.15.26357954 medRxiv
Top 2%
0.0%
Show abstract

White matter lesion (WML) identification, assessment, and characterization using magnetic resonance imaging (MRI) are fundamental for diagnosis and monitoring of multiple sclerosis (MS). Portable ultra-low field (pULF) MRI at 64 millitesla (mT) has been shown to visualize WML with at least one dimension greater than 4 mm. An automated WML segmentation tool catered to pULF-MRI can provide standardized and accurate quantitative measurements of WML volume. In this study, we sought to investigate and compare the accuracy of machine-learning (ML) and deep-learning (DL) pULF MRI segmentation tools. Same-day paired pULF (64mT) and high-field (HF, 3T) MRI scans from 84 adults with MS or suspected-MS (mean age {+/-} SD: 48 {+/-} 13, 62 females) included T2-FLAIR and T1w images. Reference WML segmentations were manually annotated on pULF T2-FLAIR for all scans, with WML confirmed with registered HF T2-FLAIR. HF reference WML segmentations were created. Four automated segmentation methods were applied to pULF scans: Method for Inter-Modal Segmentation Analysis (MIMoSA), an ML algorithm trained on HF WML masks; WMH-SynthSeg, a convolutional neural network model with flexible segmentation capabilities across field strengths and resolution; nnU-Net, a DL algorithm trained on pULF reference WML masks; and Pseudo-Label Assisted nnU-Net (PLAn), a DL algorithm pre-trained on HF reference WML masks and refined with 64mT reference WML masks. Two models were trained with nnU-Net, one using T2-FLAIR images only (nnU-Net-FL) and one using T1w and T2-FLAIR images (nnU-Net-FL/T1). The same was done with PLAn, creating PLAn-FL and PLAn-FL/T1. The six automated WML segmentation outputs were compared to the manual segmentations to determine Dice Similarity Coefficient (DSC) scores. Associations of WML volume estimates with clinical measures were investigated. DSC scores with pULF reference WML masks from PLAn-FL (DSC mean {+/-} SD: 0.50 {+/-} 0.24) outperformed MIMoSA (0.24 {+/-} 0.20, p < 0.0001), WMH-SynthSeg (0.30 {+/-} 0.18, p < 0.0001), nnU-Net-FL (0.41 {+/-} 0.24, p < 0.0001), and nnU-Net-FL/T1 (0.41 {+/-} 0.26, p = 0.0004). Worse Expanded Disability Status Scale (EDSS) and Scripps Neurologic Rating Scale (SNRS) scores were correlated with higher WML volumes in the pULF and HF reference masks. They were also correlated with WML volumes derived from WHM-SynthSeg, nnU-Net-FL, nnU-Net-FL/T1, PLAn-FL, and PLAn-FL/T1, but not MIMoSA. After adjusting for age, WHM-SynthSeg, nnU-Net FL, nnU-Net-FL/T1, PLAn-FL, and PLAn-FL/T1 had significant associations with EDSS and SNRS scores. nnU-Net and PLAn performed best in segmenting WML on pULF-MRI at 64 mT, providing accurate quantitative estimates of WML burden. Moreover, WML volumes estimated by these algorithms were associated with clinical measures of disability, underscoring their utility for reflecting clinical and radiological disease severity. Given pULF-MRI's mobility and lower cost, these findings highlight its relevance in clinical trials, particularly in involving more participants who face logistical constraints and barriers.

16
Microvascular Thrombosis and Acute Kidney Injury in COVID-19: A Systematic Review and Quantitative Analysis

Duarte, C. A.; Uscocovich, V. S. M.; Misael, I.; Duarte, P. D. A. C.; Sestito, E. B.; Da SIlva, P. N.

2026-07-17 nephrology 10.64898/2026.07.14.26357748 medRxiv
Top 2%
0.0%
Show abstract

Abstract Objective: To synthesize the available evidence on the association between SARS-CoV-2-related microvascular thrombosis and acute kidney injury (AKI), with emphasis on renal outcomes, mortality, and renal replacement therapy requirements. Methods: This systematic review followed the PRISMA 2020 statement and was prospectively registered in PROSPERO (CRD420251132701). PubMed/MEDLINE, Scopus, and Embase were searched for systematic reviews, including meta-analyses, and umbrella reviews investigating the association between SARS-CoV-2-related microvascular thrombosis and acute kidney injury. Two reviewers independently performed study selection, data extraction, and methodological quality assessment using AMSTAR-2 and ROBIS. Evidence was synthesized through a structured narrative synthesis supported by quantitative data extracted from the included reviews. Results: Six evidence syntheses evaluating kidney involvement, thrombotic events, and microvascular mechanisms in COVID-19 were included. AKI incidence was 9.2% (95%CI 4.6-13.9) among hospitalized patients and 32.6% (95%CI 8.5-56.6) among critically ill patients. In children with multisystem inflammatory syndrome associated with SARS-CoV-2, AKI incidence was 20% (95%CI 14-28). Microvascular or thrombotic events were associated with adverse renal outcomes (OR 2.14; 95%CI 1.32-3.48). AKI was associated with increased mortality (OR 4.68; 95%CI 1.06-20.70) and greater likelihood of renal replacement therapy requirement (OR 2.87; 95%CI 1.45-5.68). The certainty of evidence ranged from moderate to high for the principal outcomes. Conclusion: Current evidence supports an important association between microvascular thrombotic injury and COVID-19-associated AKI. These findings reinforce the relevance of endothelial dysfunction and thromboinflammatory pathways in kidney involvement during COVID-19 and highlight the need for early renal monitoring, risk stratification, and kidney-protective strategies in high-risk patients. Keywords: COVID-19; Acute Kidney Injury; Microvascular Thrombosis; SARS-CoV-2; Renal Replacement Therapy; Systematic Review

17
How Do Nurses Make Clinical Decisions Via Remote Reviews: A Convergent Mixed-Methods Study

Zhang, Y.; Sutherland, S.; GREENWAY, K.; Stayt, L.

2026-07-17 nursing 10.64898/2026.07.15.26357946 medRxiv
Top 2%
0.0%
Show abstract

Abstract Background: Remote clinical reviews have become an integral component of contemporary nursing practice across community and acute care settings. Nurses increasingly make autonomous clinical decisions using telephone, video, and online/digital systems, often with limited sensory information and under conditions of uncertainty. However, empirical understanding of how nurses make clinical decisions via remote reviews remains limited. Aim: To explore and understand how registered nurses (RNs) make clinical decisions about patient care via remote reviews. Methods: A convergent mixed-methods design was employed. Quantitative data (analytic quantitative sample N=53) were collected using validated questionnaires that measured decision-making processes, physician-nurse collaboration, decision-making stress, and perceived decision-making ability. Qualitative data (N=23) were generated through semi-structured interviews. Data collection took place between October 2024 and April 2025. Quantitative data were analysed using descriptive statistics, correlation, and multiple regression. Qualitative data were analysed using framework analysis. Integration was achieved through pillar-building and theory-driven synthesis and illustrated by joint display tables. Results: Most nurses demonstrated a flexible decision-making style, integrating analytical and intuitive reasoning. Both analytical and intuitive processes were positively associated with perceived decision-making ability. Physician-nurse collaboration emerged as a strong predictor of decision-making confidence, while decision-related stress was not a significant predictor. Qualitative findings identified three themes: characteristics of remote review; making adaptive decisions shaped by both internal and external constraints and enablers; and external influencing factors. The integrated findings informed a theory-informed ICE framework to illustrate how nurses make clinical decisions via remote reviews. Conclusion: Remote clinical decision-making is a dynamic cognitive-environmental process rather than a purely individual cognitive act. The ICE framework conceptualises this interaction, extending existing decision-making theories to digitally mediated care. Impact: Understanding remote decision-making supports training design, clinical governance, and the development of Artificial Intelligence-enhanced decision-support tools grounded in ecological bounded rationality. Patient or Public Contribution: Patient and public representatives contributed to stakeholder discussions that informed the development of the interview topic guide and the theoretical model. Patients or members of the public were not involved in recruitment, data collection, analysis, interpretation of findings, or preparation of the manuscript. Keywords: clinical decision-making, remote reviews, telehealth, nursing, mixed methods, ecological bounded rationality

18
Photobiomodulation promotes wound healing and functional improvement following lumbar decompression surgery: a double-blinded, placebo-controlled study

Rivera, J.; Zhou, Y.; Sak, L.; Pudewa, F.; Lee, J.; Yamamoto, M. T.; Yoo, H.; Lum, M.; Zhang, M.; Patel, A.; Vandenberghe, L. E.; Fenn, S. K.; Wang, Y.; Bailey, B.; Holley, S. M.; Vivas, A. C.; Holly, L. T.; Lu, D. C.

2026-07-17 surgery 10.64898/2026.07.15.26357882 medRxiv
Top 2%
0.0%
Show abstract

Objective: Photobiomodulation therapy has emerged as a promising modality to facilitate scar healing and pain management in dermatology and plastic surgery. However, its role in postoperative care following spine surgeries remains understudied. This double-blinded, placebo-controlled study aimed to investigate the effects of photobiomodulation in patients with chronic lower back pain undergoing lumbar decompression, with postoperative wound healing as the primary outcome and pain reduction and functional recovery as secondary outcomes. Methods: Patients were randomized to receive either active photobiomodulation braces (N=13) or placebo braces (N=12). Follow-up assessments were performed at 2, 4, 6, 8, and 12 weeks postoperatively. Outcomes included wound healing (Stony Brook Scar Evaluation Scale), back and leg pain (Visual Analog Scale), quality of life (EuroQol 5D), and functional status (Oswestry Disability Index). Results: Compared to the placebo group, the photobiomodulation treatment group had a 4.12-fold cumulative improvement in final scar scores, with significant between-group differences at postoperative weeks 6, 8, and 12 (p = 0.0062, 0.010, 0.042). Among patients with severe preoperative disability, treatment resulted in a 1.89-fold faster improvement in back pain (p=0.025) and a 1.80-fold faster improvement in ODI scores (p=0.025); and superior treatment effect on wound healing were again observed at weeks 6, 8, and 12. Among patients with poor initial scars, treatment led to a significantly better scar outcome than placebo at week 6 and a 1.94-fold faster EQ5D improvement (p=0.052), with significant gains observed as early as two weeks after surgery. There were no adverse events associated with photobiomodulation treatment. Conclusions: Photobiomodulation significantly promoted postoperative wound healing following lumbar decompression surgery, with therapeutic benefits preserved even in patients with poor baseline scar scores and functional impairment. This indicates that the efficacy of photobiomodulation is not limited by the initial scar condition or disability, supporting its broad clinical applicability. Additionally, patients with severe preoperative disability experienced greater benefits from photobiomodulation than placebo, including faster reduction in back pain and more rapid improvement in functional capacity, highlighting its role in postoperative pain management and rehabilitation. These therapeutic effects are likely mediated by photobiomodulation-induced reduction of inflammation and enhancement of tissue repair. Together, this study suggests that photobiomodulation can be a promising adjunct therapy to facilitate postoperative recovery in patients undergoing spine surgery.

19
Association between serum CEA levels and ctDNA-detected Epidermal Growth Factor Receptor mutations in lung adenocarcinoma

Roy, S.; Soroar, M. K. I.; Ara, H.; Nur, S. A.; Akanda, R. A.; Saha, S.; Alam, M. M.

2026-07-17 oncology 10.64898/2026.07.14.26358115 medRxiv
Top 2%
0.0%
Show abstract

Background with objective: Detecting EGFR mutations is critical for treating lung adenocarcinoma with highly effective targeted therapies. However, standard genetic testing is expensive, complex, and often unavailable in resource-limited settings like Bangladesh. Because elevated serum CEA has been linked to these genetic alterations, it could serve as an accessible screening tool. This study aims to evaluate the association between serum CEA levels and EGFR mutation status to determine if routine CEA testing can reliably predict these mutations and guide treatment. Methodology: In this cross-sectional analytical study, we recruited 58 patients with histologically confirmed treatment naive lung adenocarcinoma. The presence of EGFR mutations in the ctDNA was determined via ARMS (Amplification Refractory Mutation System) PCR. Patient data was statistically analyzed to assess the diagnostic correlation between serum CEA levels and the presence of EGFR mutations. Result: The overall EGFR mutation rate was 43.1% with exon 19 deletion (48%) and exon 21 mutations (44%) were the predominant types. Median serum CEA levels were significantly higher in patients with EGFR mutations compared to wild-type cases (14.6 ng/ml vs 2.8 ng/ml, p<0.001). A multivariate analysis revealed a 14% increased likelihood of an EGFR mutation for 1 ng/ml rise in serum CEA. Furthermore, serum CEA showed strong diagnostic accuracy for ctDNA samples at a 6.39 ng/ml cut-off (AUC 0.82, sensitivity 68.0%, specificity 84.8%). Conclusion: Serum CEA is a valuable, cost-effective, and non-invasive biomarker demonstrating significantly higher levels and strong diagnostic accuracy in EGFR-mutated lung adenocarcinoma compared to wild-type cases.

20
Machine learning and data-driven models for predicting post-stroke dysphagia: a systematic review and meta-analysis

Mohammadi Yazdi, S.; Motevaselian, M.; Khatami, S.; Radfar, N.; jourahmad, z.; Perez, H. A.

2026-07-17 neurology 10.64898/2026.07.15.26358113 medRxiv
Top 2%
0.0%
Show abstract

Background: Post-stroke dysphagia (PSD) contributes to aspiration, pneumonia, malnutrition, prolonged hospitalization and mortality. We evaluated the discrimination, validity and readiness of machine learning and data-driven prediction models for PSD-related outcomes. Methods: Following a prospectively registered protocol (PROSPERO CRD420261419259), we searched PubMed/MEDLINE, Embase, Web of Science Core Collection, CINAHL and CENTRAL from inception through June 7, 2026. Eligible studies developed or validated multivariable prediction models for PSD-related outcomes in adults with stroke. We used PROBAST and PROBAST+AI to assess risk of bias and applicability and TRIPOD+AI to evaluate reporting. Area under the curve (AUC) estimates were pooled on the logit scale with random-effects models. Results: Twenty-four studies were included and ten contributed to meta-analysis. Four studies predicting early or incident PSD yielded a pooled AUC of 0.94 (95% CI 0.60-0.99; I2 = 95.6%). Pooled AUCs were 0.84 (95% CI 0.71-0.92) for aspiration or penetration-aspiration and 0.89 (95% CI 0.24-1.00) for severe dysphagia. The exploratory analysis of all ten risk-prediction models produced an AUC of 0.90 (95% CI 0.80-0.95), but heterogeneity was substantial (I2 = 90.3%) and the prediction interval was 0.51-0.99. Every study had high risk of bias because of analysis-domain concerns; calibration and external validation were uncommon. Conclusions: Reported discrimination was often high, but the evidence does not establish reliable performance in care. Independent validation, calibration, complete model reporting and clinical-impact studies are needed before these models guide post-stroke swallowing care. Keywords: Post-stroke dysphagia; Stroke; Deglutition disorders; Machine learning; Clinical prediction model; Area under the curve; Meta-analysis